Papers with speaker diarization
An automated medical scribe for documenting clinical encounters (N18-5)
Copied to clipboard
Gregory Finley, Erik Edwards, Amanda Robinson, Michael Brenndoerfer, Najmeh Sadoughi, James Fone, Nico Axtmann, Mark Miller, David Suendermann-Oeft
| Challenge: | a medical scribe is a clinical professional who charts patient–physician encounters in real time. |
| Approach: | They propose to use multiple speech and language technologies to create an automated medical scribe. |
| Outcome: | a medical scribe can be used as an alternative to human scribes or as an assistive tool for physicians . the system relies on multiple speech and language technologies, including speaker diarization, medical speech recognition, knowledge extraction, and natural language generation. |
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models (2026.eacl-long)
Copied to clipboard
Iwona Christop, Mateusz Czyżnikiewicz, Paweł Skórzewski, Łukasz Bondaruk, Jakub Kubiak, Marcin Lewandowski, Marek Kubis
| Challenge: | Existing benchmarks for testing audio modality of multimodal large language models focus on testing audio tasks in isolation. |
| Approach: | They propose a new benchmark to assess multimodal large language models' ability to combine audio tasks. |
| Outcome: | The proposed benchmarks show that multimodal models can solve problems that require reasoning over audio signals with satisfactory results. |
ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | a meta corpus of audio files is used to gather, annotate and transcribe speech . a large number of speech databases are needed to perform multi-speaker tasks such as speaker diarization and speaker change detection. |
| Approach: | They propose to use human feedback to homogenize and correct speaker labels among the audio files by integrating human feedback within a speaker verification system. |
| Outcome: | The proposed protocol evaluates speech segmentation, speaker diarization, speech transcription and speaker change detection using human feedback. |
Indigenous language technologies in Canada: Assessment, challenges, and successes (C18-1)
Copied to clipboard
Patrick Littell, Anna Kazantseva, Roland Kuhn, Aidan Pine, Antti Arppe, Christopher Cox, Marie-Odile Junker
| Challenge: | There are approximately 60 Indigenous languages currently spoken in Canada. |
| Approach: | They examine which technologies have been developed and which are feasible to develop for the 60 Indigenous languages spoken in Canada. |
| Outcome: | The proposed technologies are based on the existing technologies and are feasible for most or all of these languages. |
Computer-assisted Speaker Diarization: How to Evaluate Human Corrections (L18-1)
Copied to clipboard
| Challenge: | a framework to evaluate the human corrections of a speaker diarization is presented for the French National Audiovisual Institute (INA) the speaker diaarization task is a necessary pre-processing step for speaker identification and speech transcription. |
| Approach: | They propose a framework to evaluate the human corrections of a speaker diarization . they propose four elementary actions to correct the diarized speaker and an automaton to simulate the correction sequence. |
| Outcome: | The proposed framework copes with the needs of the French National Audiovisual Institute (INA) due to the increasing number of documents and the limited number of annotators, many documents remain undocumented or only partly documented. |
A Comprehensive Evaluation of Incremental Speech Recognition and Diarization for Conversational AI (2020.coling-main)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) systems are increasingly powerful and more numerous with several options existing as a service. |
| Approach: | They evaluate the most popular automatic speech recognition systems with metrics and experiments designed with these standards in mind. |
| Outcome: | The most popular ASR systems are Microsoft and IBM, and none are suitable for natural spontaneous conversations in real-time. |
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification. (2022.lrec-1)
Copied to clipboard
Rémi Uro, David Doukhan, Albert Rilliard, Laetitia Larcher, Anissa-Claire Adgharouamane, Marie Tahon, Antoine Laurent
| Challenge: | Existing methods for creating diachronic corpus of voices are based on speaker characteristics and require human intervention. |
| Approach: | They propose to use a semi-automatic pipeline to create a diachronic corpus of voices balanced for speaker’s age, gender and recording period, according to 32 categories. |
| Outcome: | The proposed method cut down on manual annotations by ten and provides high quality speech for most of the selected excerpts. |
Towards end-2-end learning for predicting behavior codes from spoken utterances in psychotherapy conversations (2020.acl-main)
Copied to clipboard
| Challenge: | Xu and Sarikaya, 2014) proposes a framework for predicting utterance level labels directly from speech features. |
| Approach: | They propose a framework for predicting utterance level labels directly from speech features using a pretrained Speech-2-Vector encoder as bottleneck. |
| Outcome: | The proposed model outperforms state-of-the-art approaches which use transcribed text for the task of predicting psychotherapy-relevant behavior codes. |
Bazinga! A Dataset for Multi-Party Dialogues Structuring (2022.lrec-1)
Copied to clipboard
Paul Lerner, Juliette Bergoënd, Camille Guinaudeau, Hervé Bredin, Benjamin Maurice, Sharleyne Lefevre, Martin Bouteiller, Aman Berhe, Léo Galmant, Ruiqing Yin, Claude Barras
| Challenge: | a dataset of 16 TV and movie series is filled with challenging multi-party dialogues. |
| Approach: | They propose a dataset built around 16 TV and movie series with challenging multi-party dialogues. |
| Outcome: | The proposed dataset is a step towards better multi-party dialogue structuring and understanding. |
Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization (2023.findings-acl)
Copied to clipboard
| Challenge: | Current speaker diarization systems consider only acoustic information, resulting in performance degradation when encountering adverse acustic environment. |
| Approach: | They propose methods to extract speaker-related information from conversational semantics in multi-party meetings. |
| Outcome: | The proposed method improves on AISHELL-4 and AliMeeting datasets on speakers diarization and speaker-turn detection. |
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party Conversation (2025.acl-long)
Copied to clipboard
Luyao Cheng, Hui Wang, Chong Deng, Siqi Zheng, Yafeng Chen, Rongjie Huang, Qinglin Zhang, Qian Chen, Xihao Li, Wen Wang
| Challenge: | Mainstream speaker diarization systems rely only on acoustic information, making it challenging in complex aural environments. |
| Approach: | They propose a multimodal approach that integrates audio, visual, and semantic cues to enhance speaker diarization. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on multi-party conversations . it integrates audio-visual-semantic cues into the clustering process for acoustic speaker embeddings . |